Papers with multimodal cascaded cross-attention model

1 papers
Multimodal Intent Discovery from Livestream Videos (2022.findings-naacl)

Copied to clipboard

Challenge: Existing models for instructional video understanding struggle to understand abstract intents . identifying procedural intent within instructional videos is a challenging task .
Approach: They propose to extract instructional intent from software instructional livestreams by using a multimodal cascaded cross-attention model that integrates weaker and noisier video signals with more discriminative text signals.
Outcome: The proposed model improves on baseline models and compares it to existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations